Papers with shifting analysis
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing explainability methods for Large Language Models treat hidden states as static points in activation space, but they are saturated with polysemantic features. |
| Approach: | They propose a framework that shifts analysis from static activations to layer-wise geometric displacement. |
| Outcome: | The proposed framework outperforms existing explainability methods on commonsense reasoning, question answering, and toxicity detection benchmarks. |